Papers with speech recognition
Copied to clipboard
| Challenge: | Increasing interest in building multilingual foundation models for NLP and speech research has led to limited data collection for training ST systems. |
| Approach: | They propose to use Whisper to explore the behavior of multilingual speech foundation models with restricted data. |
| Outcome: | The proposed model can translate to Chinese with a single language, and it can perform transcriptions in other languages. |
Copied to clipboard
| Challenge: | Simultaneous translation is a problem that has long been considered one of the hardest problems in AI . this tutorial will provide a deep understanding of the history and the recent advances in simultaneous translation. |
| Approach: | This tutorial will examine the design and evaluation of policies for simultaneous translation . it will provide an overview of the history and recent advances in simultaneous translation. |
| Outcome: | This tutorial will examine the design and evaluation of policies for simultaneous translation . |
Copied to clipboard
| Challenge: | Introduction to deep Bayesian learning for natural language addresses the fundamentals of statistical models and neural networks. |
| Approach: | This tutorial addresses the advances in deep Bayesian learning for natural language . it focuses on advanced Bayessian models and deep models . authors present case studies and domain applications to tackle different issues . |
| Outcome: | This tutorial focuses on advanced Bayesian models and deep models for natural language . case studies and domain applications are presented to tackle different issues in deep Bayessian processing, learning and understanding. |
Copied to clipboard
| Challenge: | a proposed speech-based task-oriented dialogue system is built on a small embedded device . the system does not require internet connectivity because all components run locally on the device - a cost-effective solution . |
| Approach: | They propose a spoken-language end-to-end task-oriented dialogue system for small embedded devices such as home appliances. |
| Outcome: | The proposed system is based on a demo run offline on swiss raspberry pi . it eliminates privacy risks and eliminates server costs and latency . |
Copied to clipboard
| Challenge: | Medical dictation is one of the most common ways to document clinical encounters. |
| Approach: | They propose a machine callytranslation technique that automates post-processing tasks . they show that it outperforms conventional systems in correcting errors . |
| Outcome: | The proposed method outperforms conventional systems in many tasks while being much simpler to maintain. |
Copied to clipboard
| Challenge: | KT-Speech-Crawler is an automated dataset building tool for speech recognition. |
| Approach: | They propose an approach for automatic dataset construction for speech recognition by crawling YouTube videos. |
| Outcome: | The proposed algorithm can obtain 150 hours of transcribed speech in a day with an estimated 3.5% word error rate. |
Copied to clipboard
| Challenge: | Large-scale conversational assistants can cause errors in their modules . a machine learning system can analyze large volumes of data and isolate the source of error . |
| Approach: | They propose a machine learning system that embeds incoming request and context using pre-trained transformer models and encodes additional metadata features to output failure point predictions. |
| Outcome: | The proposed system obtains 92.2% of human performance while scaling to analyze the entire traffic in 8 different languages of a large-scale conversational assistant. |
Copied to clipboard
| Challenge: | Using RETURNN, we train and decode attention models for translation and speech recognition. |
| Approach: | They propose a layer-wise pretraining scheme for recurrent attention models and show its significant effect on deep recurrence encoder networks. |
| Outcome: | The proposed training and decoding scheme improves 1% on expected training and improves on WMT 2017 and Switchboard. |
Copied to clipboard
| Challenge: | a confidence score is a scalar quantity that measures the reliability of an automatic system. |
| Approach: | They propose to use a confidence measure to evaluate the reliability of an SLU system . they build confidence models for three different types of dialogue states . |
| Outcome: | The proposed model can be used to reject low-confidence SLU results in real-world scenarios. |
Copied to clipboard
| Challenge: | a new research platform supports spoken dialogue interaction with multiple robots . a ground robot and an aerial robot are used to perform search and rescue tasks . |
| Approach: | They propose a platform that supports spoken dialogue interaction with multiple robots . they use existing tools for speech recognition and dialogue management . |
| Outcome: | The proposed platform supports spoken dialogue interaction with multiple robots in a search and rescue scenario. |
Copied to clipboard
| Challenge: | Text normalization (TN) is an important step in conversational systems. |
| Approach: | They frame text normalization as a machine translation task and tackle it with sequence-to-sequence models. |
| Outcome: | The proposed model normalizes written text to its spoken form to facilitate speech recognition and text-to-speech synthesis. |
Copied to clipboard
| Challenge: | a recent development of spoken dialogue systems has enabled deep learning to achieve state-of-the-art performance. |
| Approach: | They propose a Python-based domain-independent, open-source toolkit for spoken dialogue systems. |
| Outcome: | The proposed toolkit extends OpenDial's Java-based architecture and provides new functions for neural dialogue state tracking and action planning. |
Copied to clipboard
| Challenge: | a new open-source web application for simultaneous speech-to-text translation is developed for Estonian . the system translates live Estonian speech into English, Russian, and Ukrainian text, and also supports English-to Estonian translation. |
| Approach: | They propose a web application that combines streaming speech recognition with a simultaneous translation model. |
| Outcome: | The proposed system outperforms existing streaming speech recognition systems in Estonian-to-English translation. |
Copied to clipboard
| Challenge: | a corpus of speech utterances collected in schools of northern italy is being used to assess the performance of students learning both English and German. |
| Approach: | a corpus of speech utterances collected in schools of northern italy is described . the corpus is going to be freely distributed to scientific community . |
| Outcome: | The corpus of speech utterances collected in schools of northern italy is a "Trentino Language Testing" in schools" the data are used to assess the performance of students learning English and German . |
Copied to clipboard
| Challenge: | Recent advances in language modeling have led to the emergence of large language models capable ofvarious natural language processing tasks. |
| Approach: | They propose a multi-instructional training approach that integrates a large language model with a speech encoder to harness the capabilities of LLMs for speech recognition and beyond. |
| Outcome: | The proposed model can be trained and aligned with a multilingual LLM on 1900 hours of transcribed data from 139 languages. |
Copied to clipboard
| Challenge: | Named entity recognition (NER) tasks require large labeled datasets to perform . compared to prior work, relative improvements in F1 of up to 16% are found . |
| Approach: | They propose to use self-training, knowledge distillation, and transfer learning to learn SLU models . they compare pipeline and pipeline approaches to find out how to use external data . |
| Outcome: | The proposed models improve performance beyond pre-trained models in resource-constrained settings . the best baseline model is a pipeline approach, while the best performance is achieved by an E2E model. |
Copied to clipboard
| Challenge: | Unvoiced electromyography (EMG) is an effective communication tool for individuals unable to produce vocal speech. |
| Approach: | They propose an EMG adaptor module that maps EMG features to an LLM's input space and achieves an average word error rate of 0.49 on a closed-vocabulary unvoiced EMG-to-text task. |
| Outcome: | The proposed module achieves an average word error rate of 0.49 on a closed-vocabulary unvoiced EMG-to-text task. |
Copied to clipboard
| Challenge: | The Kazakh speech corpus contains over 153,000 utterances spoken by participants from different regions and age groups, as well as both genders. |
| Approach: | They propose to build an open-source Kazakh speech corpus for the Kazakh language that contains over 153,000 transcribed audio . they describe the data collection and preprocessing procedures followed by a description of the database specifications. |
| Outcome: | The Kazakh speech corpus contains over 153,000 utterances spoken by participants from different regions and age groups, as well as both genders. |
Copied to clipboard
| Challenge: | Existing TT processes face challenges such as incomplete data collection, communication barriers, and manual errors, leading to high over-triage and under-triages rates. |
| Approach: | They propose to use an AI-driven multilingual TT system to provide decision support for triage. |
| Outcome: | The proposed system achieves word error rate of 14.57% for speech recognition and an F1 score of 73.34% for key information extraction. |
Copied to clipboard
| Challenge: | Recurrent neural networks have achieved state-of-the-art results in many artificial intelligence tasks, such as language modeling, neural machine translation and speech recognition. |
| Approach: | They propose an efficient architecture to improve the efficiency of such RNN model training by adopting the group strategy for recurrent layers while exploiting the representation rearrangement strategy between layers as well as time steps. |
| Outcome: | The proposed architecture achieves comparable or better accuracy compared with baselines, with a much smaller number of parameters and at a lower computational cost. |
Copied to clipboard
| Challenge: | Sustainable speech recognition systems are essential for scientists, journalists, and anyone processing audio recordings of interviews and meetings. |
| Approach: | They propose a speech-to-text system "Pisets" which is based on a three-component architecture aimed at improving speech recognition accuracy while minimizing errors and hallucinations associated with the Whisper model. |
| Outcome: | The proposed system ensures robust transcribing of long audio data across various acoustic conditions compared to WhisperX and the usual Whisper model. |
Copied to clipboard
| Challenge: | In this paper we explore the use of Learning Hidden Unit Contribution for neural machine translation. |
| Approach: | They propose to use Learning Hidden Unit Contribution for the task of neural machine translation. |
| Outcome: | The proposed method achieves improvements of up to 2.6 BLEU points over a general system . it also achieves up to 6 BLUE points if the initial system has been trained on out-of-domain data . |
Copied to clipboard
| Challenge: | VoxPopuli provides 400K hours of unlabeled speech data in 23 languages . large amounts of multilingual audio data are needed to achieve similar progress for multilingual ASR and ST. |
| Approach: | They propose a large-scale multilingual corpus that provides 400K hours of unlabeled speech data in 23 languages. |
| Outcome: | The proposed corpus provides 400K hours of unlabeled speech data in 23 languages and 1.8K hours transcribed speeches in 15 languages and their aligned oral interpretations into 15 target languages totaling 17.3K hours. |
Copied to clipboard
| Challenge: | Recent work in deep fusion models has led to substantial improvements over unimodal approaches in areas like speech recognition, emotion recognition and analysis. |
| Approach: | They propose to introduce neural dependencies into the loss functions to allow for fusion of different modalities while keeping the model complexity manageable. |
| Outcome: | Experiments on multimodal sentiment analysis tasks show that the proposed approach provides a consistent performance boost. |
Copied to clipboard
| Challenge: | Existing methods to pre-train speech and text use unlabeled data to learn universal feature representations. |
| Approach: | They propose a method to jointly pre-train speech and text in an encoder-decoder modeling framework for speech translation and recognition. |
| Outcome: | The proposed method achieves between 1.7 and 2.3 BLEU improvement above the state of the art on the MuST-C speech translation dataset and comparable WERs to wav2vec 2.0 on the Librispeech speech recognition task. |
Copied to clipboard
| Challenge: | Bemba is the most populous language of Zambia but lacks resources for research . despite its significance, Bemba remains under-resourced and lacking in high-quality data and resources for NLP experiments and language technologies. |
| Approach: | They propose a large multimodal dataset for Bemba that includes images, transcriptions and translations. |
| Outcome: | The proposed dataset is based on images, transcriptions and translations of Bemba speakers . it provides baselines on speech recognition, machine translation and speech translation tasks . |
Copied to clipboard
| Challenge: | Existing work has extended recurrent neural networks to model lattice inputs but these models suffer from slow computation speeds. |
| Approach: | They propose to extend the paradigm of self-attention to handle lattice inputs by adding probabilistic reachability masks that incorporate latticae structure into the model and support lattics if available. |
| Outcome: | The proposed model outperforms baseline models while being much faster to compute than previous models. |
Copied to clipboard
| Challenge: | Recent advances in multimodal and speech-native large language models have delivered impressive speech recognition, translation, understanding, and question-answering capabilities for high-resource languages. |
| Approach: | They propose to benchmark African languages and African-accented French, Arabic, and 100+ African English accents across 20 African languages. |
| Outcome: | The proposed model outperforms traditional speech transcription and translation models in African languages and non-native French or English accents. |
Copied to clipboard
| Challenge: | AccentFold uses spatial relationships to improve speech recognition for accented speech . existing methods for accent recognition have been limited due to data scarcity and budget constraints . |
| Approach: | They propose a method that exploits spatial relationships between learned accent embeddings to improve downstream automatic speech recognition. |
| Outcome: | The proposed method outperforms baseline methods in accented speech training. |
Copied to clipboard
| Challenge: | Conventionally, neural language models are trained by minimizing perplexity (PPL) on grammatical sentences. |
| Approach: | They propose a large margin criterion for training neural language models by minimizing perplexity on grammatical sentences and propose enlarged margins for task-specific training. |
| Outcome: | The proposed method gains up to 1.1 WER reduction for speech recognition and 1.0 BLEU increase for machine translation. |
Copied to clipboard
| Challenge: | Existing methods to regularize multimodal data are imperfect due to imperfect modalities, missing entries or noise corruption. |
| Approach: | They propose a method to regularize multimodal data by tensor rank minimization . they use correlations between time and modalities to generate low-rank tenses . |
| Outcome: | The proposed model achieves good results across various levels of imperfection. |
Copied to clipboard
| Challenge: | End-to-end speech translation (ST) models need large amount of training data to perform well. |
| Approach: | They propose a shrinking mechanism to mitigate the length mismatch between speech and text features by predicting word boundaries. |
| Outcome: | The proposed method achieves better performance on the MUST-C dataset, with higher inference speed and lower memory usage. |
Copied to clipboard
| Challenge: | Currently, partitioning speech corpora is done by hand, but this is not feasible for the dataset under investigation. |
| Approach: | They propose to partition a 41.6-hour corpus of code-switched speech into training, development and testing partitions using mixed-integer linear programming. |
| Outcome: | The proposed method allows to partition a 41.6-hour corpus of code-switched speech into training, development and testing partitions while maintaining a fixed number of speakers and a specific amount of codeswitching speech in the development and test partitions. |
Copied to clipboard
| Challenge: | despite its crucial role in research experiments, code correctness is often presumed on the perceived quality of results. |
| Approach: | They propose to promote code-quality checklists to promote coding best practices . they propose to fix bugs in conformer implementations to mitigate this risk . |
| Outcome: | The proposed checklists aim to promote coding best practices and improve software quality within the NLP community. |
Copied to clipboard
| Challenge: | Autoregressive coding targets are used to learn meaningful representations from unlabeled speech. |
| Approach: | They propose a method that trains an autoregressive RNN to generate an unseen future frame given a context such as recent past frames. |
| Outcome: | The proposed method can learn representations from unlabeled speech. |
Copied to clipboard
| Challenge: | x-vector (speaker recognition PTM) achieves the highest performance in prosodic tasks . despite its low parameter, x vector captures unique prosodic characteristics of the sources . |
| Approach: | They propose to use SOTA speech pre-trained models to capture prosodic sig-natures of generative sources for audio deepfake source attribution. |
| Outcome: | The proposed model captures prosodic sig-natures of generative sources better than other models on ASVSpoof and CFAD. |
Copied to clipboard
| Challenge: | There are approximately 60 Indigenous languages currently spoken in Canada. |
| Approach: | They examine which technologies have been developed and which are feasible to develop for the 60 Indigenous languages spoken in Canada. |
| Outcome: | The proposed technologies are based on the existing technologies and are feasible for most or all of these languages. |
Copied to clipboard
| Challenge: | Existing models rely on global visual features that represent the entire image, but localizing the relevant regions of the image will make it possible to recover a larger set of words, such as adjectives and verbs. |
| Approach: | They propose a multimodal automatic speech recognition system that uses visual information from different parts of the image to ground the speech in the visual context. |
| Outcome: | The proposed model improves over approaches that use global visual features and localizes the correct proposals. |
Copied to clipboard
| Challenge: | Samrómur is the largest prompted speech collection effort for Icelandic so far and verification is as monumental as the collection itself. |
| Approach: | They propose to collect large and diverse corpus for automatic speech recognition and similar tools using crowd-sourced donations. |
| Outcome: | The collected utterances are based on the Mozilla Common Voice platform and are available for free on the Samrómur collection platform. |
Copied to clipboard
| Challenge: | Weighted finite state transducers (FSTs) are used in language processing . a GPU implementation of the composition operation is currently under development . |
| Approach: | They propose a GPU implementation of the composition operation for weighted finite state transducers. |
| Outcome: | The proposed approach achieves speedups of up to 6 times over the serial implementation and 4.5 times over OpenFST on the GPU. |
Copied to clipboard
| Challenge: | Text and vision foundation models can perform many tasks in a zero-shot setting . however, there has been little work on the zero-shoot abilities of ASR foundation models . |
| Approach: | They investigate the ability of ASR foundation models to perform zero-shot audio classification using text prompts and a decoding probability generator. |
| Outcome: | The proposed model outperforms state-of-the-art models on audio classification datasets without training them on extra data or adding any parameters. |
Copied to clipboard
| Challenge: | Existing approaches to metric meta-evaluation focus on general statements about absolute and relative quality of metrics across arbitrary system outputs, but in practice, metrics are applied in highly contextual settings. |
| Approach: | They propose a method for contextual metric meta-evaluation by comparing local metric accuracy. |
| Outcome: | The proposed method compares the local metric accuracy of evaluation metrics across translation, speech recognition, and ranking tasks. |
Copied to clipboard
| Challenge: | Disfluency removal is an intermediate step between speech recognition and machine translation (MT) with the rise of end-to-end speech translation systems, disfluency recognition and removal needs to be incorporated into the model architectures or handled as a post-processing step. |
| Approach: | They propose to use a sequence-to-sequence model to translate from noisy, disfluent speech to fluent text with disfluencies removed using the recently collected ‘copy-edited’ references for the Fisher Spanish-English dataset. |
| Outcome: | The proposed model generates fluent translations from disfluent speech using the recently collected ‘copy-edited’ references for the Fisher Spanish-English dataset. |
Copied to clipboard
| Challenge: | Currently, Transformer-based text classifiers are not suitable for live incremental processing, operating only on the level of complete sentence inputs. |
| Approach: | They propose to introduce a method for word-by-word left-to-right incremental processing to Transformers such as BERT, models without an intrinsic sense of linear order. |
| Outcome: | The proposed method maintains high non-incremental performance while operating strictly incrementally. |
Copied to clipboard
| Challenge: | Existing interpretability methods have been proposed to interpret the inner workings of Transformer models at different levels of precision and complexity. |
| Approach: | They propose a method to analyze encoder-decoder Transformers by using the decoder module Model Output encoder to cross-attend representations of intermediate encoder activations instead of using the default output. |
| Outcome: | The proposed method maps uninterpretable representations to human-interpreted sequences of words or symbols, shedding new light on the information flow in this popular but understudied class of models. |
Copied to clipboard
| Challenge: | Recent studies have augmented large language models (LLMs) with speech capabilities, leading to the development of speech language models. |
| Approach: | They propose a single-stage joint speech-text SFT approach for training SpeechLMs . their model combines text-only SFT data with three types of speech-related data . |
| Outcome: | The proposed model outperforms previous SpeechLMs on speech-based QA tasks while maintaining original speech-only capabilities. |
Copied to clipboard
| Challenge: | End-to-end speech translation requires a powerful encoder to transcribe, understand and learn cross-lingual semantics simultaneously. |
| Approach: | They propose a curriculum pre-training method that includes an elementary course for transcription learning and two advanced courses for understanding the utterance and mapping words in two languages. |
| Outcome: | The proposed method improves on En-De and En-Fr speech translation benchmarks. |
Copied to clipboard
| Challenge: | Increasing number of people in the world today speak a mixed-language as a result of being multilingual. |
| Approach: | They propose a method to transfer learn on a code-switched speech recognition system by extracting information from high-resource monolingual datasets. |
| Outcome: | The proposed model outperforms baselines on speech recognition and language modeling tasks and is faster to converge. |
Copied to clipboard
| Challenge: | Low-resource languages still lag behind in documenting endangered languages . a large corpus of culturally significant conversations is available for computational experiments . |
| Approach: | They propose a resource for computational experiments on Mapudungun, a polysynthetic indigenous language spoken in Chile. |
| Outcome: | The proposed corpus provides 142 hours of culturally significant conversations in Mapudungun . the language is spoken by the Mapuche people of southern Chile and western argentina . |
Copied to clipboard
| Challenge: | Recent research has focused on low-resource languages and the transcription bottleneck paradigm. |
| Approach: | They propose to use a spoken term detection system to train a speech recognition system in an Aboriginal community to reach better comprehension and engagement from Aboriginal participants. |
| Outcome: | The proposed system can be implemented in an Aboriginal community and reach better comprehension and engagement from Aboriginal participants. |
Copied to clipboard
| Challenge: | a new national language technology programme for Icelandic is described . the programme aims to make Icelandic usable in communication and interactions in the digital world . |
| Approach: | They describe a new national language technology programme for Icelandic . the programme aims to make Icelandic usable in communication and interactions in the digital world . |
| Outcome: | The proposed programme aims to make Icelandic usable in communication and interactions in the digital world. |
Copied to clipboard
| Challenge: | Existing studies have shown that large language generation models disadvantaging African American Language (AAL) can be biased for certain language varieties, but there is little research on the impact of these biases on other languages. |
| Approach: | They evaluate how well LLMs understand African American Language (AAL) in comparison to white Mainstream English (WME) using a dataset of AAL texts from a variety of regions and contexts, they find dialectal bias in six pre-trained LLM. |
| Outcome: | The proposed models understand African American language in comparison to white mainstream English (WME) the proposed models have performance gaps on two tasks that are not matched by the model. |
Copied to clipboard
| Challenge: | Existing approaches to reduce label noise rely on heuristics and sample losses. |
| Approach: | They propose a method that transfers the noise distribution to a clean set and trains a model to distinguish noisy labels from clean ones using model-based features. |
| Outcome: | Empirically, the proposed approach improves over strong baselines on a wide range of tasks including text classification and speech recognition. |
Copied to clipboard
| Challenge: | a corpus of sentence-aligned triples of German audio, German text, and English translation is available for speech recognition . a large corpus is available to date for end-to-end speech translation based on parallel data . |
| Approach: | They present a corpus of sentence-aligned triples of German audio, German text, and English translation based on German audio books. |
| Outcome: | The proposed corpus is the largest resource for German speech recognition and for end-to-end German-to English speech translation. |
Copied to clipboard
| Challenge: | mental health care is a demanding occupation, resulting in a severe gap in patient-centered care . a recent study shows that natural language processing can extract certain aspects of human-human communication. |
| Approach: | They propose to use data from psychotherapy sessions to help improve quality of care . they use feedback and cooperation annotations to assess quality of therapy sessions . |
| Outcome: | The proposed method aims to analyse psychotherapy data and assess its quality . it aims at identifying what qualifies for good feedback or cooperation in therapy sessions . |
Copied to clipboard
| Challenge: | Spoken language understanding (SLU) tasks have received little attention and resources compared to lower-level tasks like speech and speaker recognition. |
| Approach: | They propose annotated SLU benchmark tasks based on freely available speech data to complement existing benchmarks and address gaps in the evaluation landscape. |
| Outcome: | The proposed benchmarks complement existing benchmarks and address gaps in the evaluation landscape. |
Copied to clipboard
| Challenge: | Existing methods for speech recognition suffer from the synthetic-to-real gap . existing methods suffer from this distributional shift due to acoustic mismatches . |
| Approach: | They propose to use task arithmetic to fine-tune an ASR model on synthetic data to mitigate the synthetic-to-real gap. |
| Outcome: | The proposed method shows an improvement of 10.03% over baselines on the SLURP dataset. |
Copied to clipboard
| Challenge: | Common Voice is a massively-multilingual collection of transcribed speech intended for speech technology research and development. |
| Approach: | They propose to use Mozilla’s DeepSpeech Speech-to-Text toolkit to perform multilingual automatic speech recognition experiments. |
| Outcome: | The proposed corpus is the largest in the public domain for speech recognition, both in terms of hours and languages. |
Copied to clipboard
| Challenge: | the Huqariq corpus is a multilingual collection of speech from native Peruvian languages . the project is designed to preserve endangered languages in the public domain . |
| Approach: | They propose to use crowdsourcing to collect transcribed audio from native Peruvian languages . they propose to do 220 hours of speech recognition experiments to verify quality . |
| Outcome: | The Huqariq corpus is a multilingual collection of speech from native Peruvian languages . the project is expected to reach 20 native languages out of 48 native languages by 2022 . |
Copied to clipboard
| Challenge: | Compounding is a common word-formation process in Germanic languages . high productivity and low corpus frequency of compounds increase vocabulary size . |
| Approach: | They develop a deep learning-based approach to noun compound splitting and idiomatic compound detection for the German language. |
| Outcome: | The proposed approach outperforms the current state of the art in noun compound splitting and idiomatic compound detection for the German language. |
Copied to clipboard
| Challenge: | Named entity recognition is usually made through a pipeline process that consists of processing audio and applying a NER to the audio outputs. |
| Approach: | They propose an original 3-pass approach and explore the capability of an E2E system to do structured NER. |
| Outcome: | The proposed system performs better than the current pipeline approach. |
Copied to clipboard
| Challenge: | a systematic study on multilingual and cross-lingual intent detection from spoken data is presented . current work on intent detection is limited to English, and standard benchmarks exist only in English. |
| Approach: | They present a systematic study on multilingual and cross-lingual intent detection from spoken data. |
| Outcome: | The proposed resource is called MInDS-14, and it provides strong intent detection in most target languages. |
Copied to clipboard
| Challenge: | despite of the great demand, there is still a huge shortage in available corpora for dialectal languages and code-switched speech. |
| Approach: | They collect conversational Egyptian Arabic spontaneous speech, extract transcriptions and analyze it from a code-switching perspective. |
| Outcome: | The authors collect conversational Egyptian Arabic spontaneous speech, extract transcriptions and analyze speech from the code-switching perspective. |
Copied to clipboard
| Challenge: | Recent advances in sequence modeling have highlighted the strengths of the transformer architecture. |
| Approach: | They propose a general lattice transformer for speech translation where the input is the output of the automatic speech recognition (ASR) they propose 'controllable' lattica attention mechanism to consume latent representations. |
| Outcome: | The proposed model outperforms baseline and lattice LSTM on the Chinese-English translation task. |
Copied to clipboard
| Challenge: | Until recently, the only feasible approach to translating acoustic speech signals into text was the cascaded approach. |
| Approach: | They propose a classification of the main challenges of traditional approaches to speech translation . they argue that end-to-end models fall short due to compromises made to address data scarcity . |
| Outcome: | This paper provides a brief survey of the main challenges of traditional approaches in speech translation . it reveals that many end-to-end models fail due to compromises made to address data scarcity. |
Copied to clipboard
| Challenge: | Terminology is also needed in AI applications such as machine translation, speech recognition, information extraction, and other natural language processing tools. |
| Approach: | They propose a terminology management solution that facilitates standards-based sharing and management of terminology resources by providing the EuroTermBank Toolkit. |
| Outcome: | The EuroTermBank Toolkit facilitates standards-based sharing and management of terminology resources by participating in federated databases. |
Copied to clipboard
| Challenge: | Pre-trained speech models have advanced speech-related tasks, including speech recognition and translation. |
| Approach: | They propose a pre-trained speech model that incorporates modifications to ensure consistent speech representations during training and inference phases for streaming speech inputs. |
| Outcome: | The proposed model outperforms baseline models on speech recognition and translation tasks and achieves a superior balance between quality and latency. |
Copied to clipboard
| Challenge: | The video annotation for speech technologies corpus contains 2900 hours of video data . the data are intended to support speech technology development . |
| Approach: | The Video Annotation for Speech Technologies corpus contains 2900 hours of video data . the data are intended to support speech technology development . |
| Outcome: | The video annotation for speech technologies corpus contains 2900 hours of video data . the data are intended to support speech detection, language identification, speaker identification, and speech recognition . |
Copied to clipboard
| Challenge: | In-car smart assistants should be able to process general as well as car-related commands and perform corresponding actions, which eases driving and improves safety. |
| Approach: | They propose a dataset for in-car command recognition in the cantonese language with both video and audio data. |
| Outcome: | The proposed model can achieve a considerable quality on the clean test set, but the speech recognition quality on noisy data is still inferior. |
Copied to clipboard
| Challenge: | a recent study evaluated off-the-shelf automatic speech recognition systems . current state-of-the art systems perform poorly in domains that require special vocabulary and language models . |
| Approach: | They evaluate off-the-shelf automatic speech recognition systems across different dialogue domains . they use data collected from deployed spoken dialogue systems and human-human conversations . |
| Outcome: | The evaluation is aimed at non-experts with limited experience in speech recognition . the results show that the performance of each speech recognizer can vary significantly depending on the domain . |
Copied to clipboard
| Challenge: | transcribed-like data is often used to correct recurring errors, but training with synthetic data is difficult. |
| Approach: | They propose to use synthetic transcribed-like data to train error correction models . they show that synthetic data outperforms the common approach of random perturbations . |
| Outcome: | The proposed method outperforms the common method using random perturbations in transcribed data and language-specific adjustments to the vocabulary of a BPE tokenizer. |
Copied to clipboard
| Challenge: | In an aging society, a highly accurate speech recognition system is needed for use in electronic devices for the elderly but this cannot be achieved using conventional speech recognition systems due to the unique features of the speech of elderly people. |
| Approach: | They construct a new corpus of elderly Japanese speech from existing Japanese speech corpora and train them using existing data. |
| Outcome: | The proposed models achieve word error rates (WER) as low as 13.38%, exceeding the results of the previous study. |
Copied to clipboard
| Challenge: | Existing data sets in Hungarian are limited in quality and quality . however, it is difficult to train a modern automatic speech recognition system with thousands of hours of transcribed speech. |
| Approach: | They propose to analyze available speech data sets in Hungarian in five categories . they estimate that the available data sets are 2800 hours across 7500 speakers . |
| Outcome: | The available data sets in spoken Hungarian are compared to other languages and are estimated to be 2800 hours in size . however, their distribution and alignment to real-life tasks are far from optimal indicating the need for larger-scale natural language speech data sets. |
Copied to clipboard
| Challenge: | Existing models that incorporate audio-related image information do not improve speech recognition performance. |
| Approach: | They propose a novel approach utilizing audio-related image information and set up a multimodal speech recognition system that uses vision as hotwords to enhance the model’s speech recognition capability. |
| Outcome: | The proposed model outperforms unimodal ASR model and achieves SOTA among existing image-based multimodal ASL models. |
Copied to clipboard
| Challenge: | Existing efforts to improve robustness of audio-visual speech recognition with visual information focus on audio modality . current approaches introduce noise adaptation techniques to improve reliability of AVSR task . |
| Approach: | They propose a visual-invariant modality to strengthen robustness of audio-visual speech recognition (AVSR) it can adapt to any testing noises without dependence on noisy training data, a.k.a., unsupervised noise adaptation. |
| Outcome: | The proposed method outperforms existing state-of-the-arts on visual speech recognition task under various noisy and clean conditions. |
Copied to clipboard
| Challenge: | Recent Large Language Model (LLM) based AVSR systems incur high computational costs due to high temporal resolution of audio-visual speech. |
| Approach: | They propose an efficient multimodal speech LLM framework that minimizes token length while preserving essential linguistic content. |
| Outcome: | The proposed approach reduces token usage by 86% while using only 3.5 tokens per second. |
Copied to clipboard
| Challenge: | BLSP-Emo model understands both semantics and emotions in speech and generates empathetic responses. |
| Approach: | They propose a language-speech pretraining with emotion support that utilizes existing speech and emotion recognition datasets to create an end-to-end speech-language model. |
| Outcome: | The proposed model can understand both semantics and emotions in speech and generate empathetic responses. |
Copied to clipboard
| Challenge: | Autoregressive speech token generation models suffer from hallucinations and undesired vocalizations that do not conform to conditioning inputs. |
| Approach: | They propose an encoder-decoder transformer model that improves contextual adherence of speech token generation LLMs through preference alignment and classifier-free guidance. |
| Outcome: | The proposed model outperforms previous LLM-based models on intelligibility, speaker similarity and naturalness. |
Copied to clipboard
| Challenge: | Existing methods for estimating speech recognition metrics depend on ground truth labels. |
| Approach: | They propose a label-free approach to approximating ASR performance metrics . they embed multimodal embeddings in a unified space for speech and transcription representations . |
| Outcome: | The proposed method outperforms baseline models on speech recognition benchmarks by 50%. |
Copied to clipboard
| Challenge: | Podcasts and other audiovisual content are becoming more and more a part of everyday communication and the digital age is changing from text to voice. |
| Approach: | They synthesize the current state of the field and highlight the need for realistic evaluation benchmarks and multilingual datasets. |
| Outcome: | The proposed frameworks are based on evaluation protocols and datasets and highlight the need for realistic benchmarks and multilingual datasets. |
Copied to clipboard
| Challenge: | Similar to humans, animals make extensive use of verbal and non-verbal forms of communication, including audio signals. |
| Approach: | They propose to use self-supervised speech representation models pre-trained on human speech to address dog bark classification tasks. |
| Outcome: | The proposed model improves dog recognition, breed identification, gender classification, and context grounding tasks. |
Copied to clipboard
| Challenge: | Currently, there are no publicly available speech recognition datasets in the medical domain due to privacy restrictions. |
| Approach: | They present a Vietnamese speech recognition dataset in the medical domain comprising 16h of labeled medical speech, 1000h of unlabeled medical and 1200h of general-domain speech. |
| Outcome: | The proposed model outperforms state-of-the-art models from 51.8% to 29.6% WER on test set. |
Copied to clipboard
| Challenge: | naively fine-tuning an omni-model on speech recognition and external sound understanding tasks often degrades performance . Xie and Wu's framework, Speech-Hands, recasts the problem as an explicit self-reflection decision. |
| Approach: | They propose a voice-agentic framework that learns one critical omni-understanding skill: trusting itself versus external audio perception. |
| Outcome: | The proposed framework outperforms baseline models on the OpenASR leaderboard by 12.1% WER and high F1 on audio QA decisions. |